Phase 2: Data & Mathematics Lesson 2 of 5

Statistics
Intuition for AI

AI models are, at their core, very sophisticated statistics engines. You do not need to derive any formulas. But you do need to understand what statistics is actually telling you, and why that matters enormously for building AI that works.

You will learn
Mean, median and mode, and when each one matters
Variance and standard deviation as measures of spread
How distributions shape AI behaviour
Correlation and why it is not causation

Why does statistics matter for AI?

Before you write your first machine learning model, here is something important to understand: every model you train is making statistical decisions. When a model decides whether an email is spam, it is essentially asking, "Based on the patterns in thousands of past emails, what is the most likely category for this one?" That is a statistical question.

When a model predicts a house price, it is finding the statistical relationship between features like square footage, location and number of bedrooms and the prices of similar houses sold in the past. Statistics is not a separate subject from AI. It is the language AI speaks.

The good news is that for an intuitive understanding of AI, you do not need to be able to derive these formulas from scratch. You need to understand what they are telling you about your data.

"Statistics is the grammar of science."

Karl Pearson, mathematician

Measuring the centre: mean, median and mode

The first question you ask about any dataset is: what is a typical value? There are three ways to answer that, and they can give you very different answers.

Consider the ages of passengers on the Titanic: 22, 38, 26, 35, 35, 31, 54, 2, 27, 14.

Mean (average)
28.4
Add all values and divide by the count. Affected by extreme values (outliers). The millionaire who joins a room raises the average wealth of everyone in it.
Median (middle)
29
The middle value when sorted. Resistant to outliers. House price reports use median rather than mean because a few mansions would completely distort the average.
Mode (most common)
35
The value that appears most often. Most useful for categorical data. What is the most common class of ticket? The most common country of origin?
This is where AI bias can start

If your training data is skewed toward one group, the mean will be pulled toward that group. An AI model trained on this data will then perform better for that group and worse for others. Understanding measures of centre is not just a statistics exercise. It is a bias detection skill.

Measuring spread: variance and standard deviation

Knowing the average is not enough. You also need to know how spread out the data is. Consider two test score results where both classes had an average score of 70:

Test scores: both classes average 70, but with very different distributions
Class A: scores cluster tightly around 70. Class B: scores range widely from 35 to 97. Both classes average exactly 70, yet they are completely different in character. A model that only sees the average would treat them identically, which would be a serious mistake.

Variance measures how far, on average, each data point is from the mean. A high variance means the data is spread out. A low variance means values are clustered close to the mean.

Standard deviation is simply the square root of variance. It puts the spread back in the same units as the original data, making it easier to interpret. If house prices have a mean of £300,000 and a standard deviation of £80,000, you know that most houses fall between about £220,000 and £380,000.

Think of it this way

Imagine two archers. Both hit the target in the same average position, dead centre. But one archer's arrows are all clustered in a tight group. The other's are scattered randomly across the board. Mean tells you where the arrows land on average. Standard deviation tells you how consistent the archer is. You want both pieces of information.

Distributions: the shape of your data

A distribution shows how values are spread across the range of a dataset. The shape of that distribution tells you a lot about what is happening, and it has a direct effect on which AI algorithms will work well.

The most famous distribution is the normal distribution, also called the bell curve. Many natural phenomena follow this pattern: human heights, IQ scores, measurement errors, the weight of objects from a production line. Most values cluster around the middle. Extreme values are rare at both ends.

μ (mean) −1σ +1σ

The normal distribution (bell curve). About 68% of values fall within one standard deviation (σ) of the mean. About 95% fall within two standard deviations. This rule is called the 68-95-99.7 rule and it appears constantly in AI and statistics.

Not all data is normally distributed. Income is right-skewed, meaning most people earn moderate amounts while a small number earn enormous sums, pulling the tail to the right. Age at retirement is left-skewed. Social media engagement is heavily right-skewed: most posts get few views, a tiny number get millions. Knowing the shape of your data's distribution helps you choose the right model and understand its limitations.

Correlation: how features relate to each other

One of the most powerful ideas in statistics for AI is correlation. Two variables are correlated when they tend to move together. As one goes up, the other tends to go up (positive correlation) or down (negative correlation).

Study hours vs exam score: a positive correlation
Hours studied Exam score outlier

Each dot is one student. As hours studied increases, exam scores tend to rise, showing a clear positive correlation. The red dot is an outlier. Notice how even with an outlier, the overall trend is still visible. The dashed line is a regression line. You will build one of these yourself in Lesson 3.3.

Correlation is measured from −1 to +1. A value of +1 means perfect positive correlation. A value of −1 means perfect negative correlation (as one rises, the other falls reliably). A value near 0 means no linear relationship.

The most important warning in all of statistics

Correlation does not imply causation. Ice cream sales and drowning rates are positively correlated. Both rise in summer. That does not mean eating ice cream causes drowning. A third variable (hot weather) drives both. AI models can find and exploit correlations powerfully, but they have no concept of cause and effect unless you explicitly build that understanding in. This is one reason why AI systems can be confidently wrong.

Probability: the engine underneath

Every prediction an AI model makes is really a probability statement. A spam classifier does not say "this email IS spam." It says "there is a 94% probability this email is spam." An image classifier does not say "this IS a cat." It says "this image is 87% likely to be a cat, 9% likely to be a dog, 4% something else."

Probability runs from 0 (impossible) to 1 (certain). A probability of 0.5 means the model genuinely does not know. It is a coin flip. When you evaluate AI models, you will often set a threshold: if probability is above 0.5, classify as positive. Changing that threshold changes which errors you make more often, a trade-off you will explore in Lesson 3.6 on evaluation metrics.

Understanding that AI outputs are probabilities, not certainties, is one of the most important conceptual shifts for any beginner. It changes how you use AI, how much you trust it, and how you build systems around it.

Lesson Activity · No code required
Spot the Misleading Stat
Statistics can be used to tell accurate truths or misleading half-truths. This exercise builds the critical eye you will need to evaluate AI model results honestly.
01 Find one statistic in the news this week. Look for any headline that quotes a number, a percentage, or a "study shows..." claim.
02 Ask three questions about it: What is the sample size? Is this a mean or median? Could a third variable explain the finding? Does the statistic actually support the headline?
03 Write two sentences: what the statistic actually shows, and what conclusion would be a mistake to draw from it. Bring it to the live session.
Your Notes
Studying independently? Write your thoughts or answers below. Notes save automatically to your browser.

Reflect

Before you move on

No right answers here. These questions are for you.

A hospital reports that the "average patient wait time is 12 minutes." Why might that number mislead you, and what would you ask to see instead?

The average collapses the full distribution into one number. A handful of patients waiting 4 hours can coexist with a median of 8 minutes. In AI, models trained to minimise average error can still perform terribly for edge cases or minority groups. Ask for the distribution, not just the mean.

A coin lands heads 7 times in a row. You feel certain the next flip must be tails. An AI trained on this logic would do what, and why is that a problem?

This is the gambler's fallacy. Each flip is independent; past outcomes do not change future probabilities. An AI that falls into this trap is confusing correlation in a sample with a causal rule. Real probability reasoning requires knowing whether events are truly independent before drawing conclusions.

Cities with more cinemas tend to have higher crime rates. Should the government build fewer cinemas to reduce crime? What is the statistical concept at play?

Both cinema counts and crime rates are driven by a third variable: population size. Larger cities have more of both. This is a classic case of correlation without causation, compounded by a lurking variable. Acting on it would waste resources and miss the real factors driving crime.

Progress
Done with this lesson?
Mark it complete to track your progress.